Papers with Slavic languages

8 papers
Dialect Clustering with Character-Based Metrics: in Search of the Boundary of Language and Dialect (2020.lrec-1)

Copied to clipboard

Challenge: 'A language is a dialect with an army and navy' is attributed to sociologist Max Weinrich.
Approach: They propose a universal character-based method for representing sentences so that one can calculate the distance between any two sentence pairs.
Outcome: The proposed method can be used to calculate distance between two sentences by clustering a dialect/sub-language mixed corpus into sub-groups and to partially answer the question of what separates languages from dialects.
BERT-like Models for Slavic Morpheme Segmentation (2025.acl-long)

Copied to clipboard

Challenge: Existing morpheme segmentation algorithms for Slavic languages have been improved but performance is still low for words with roots not present in training data.
Approach: They propose to fine-tune BERT-like models for morpheme segmentation using data from Belarusian, Czech, and Russian to account for word semantics.
Outcome: The proposed models outperform all previous approaches in Czech and Russian, with word-level accuracy of 92.5-95.1%.
Cross-lingual Named Entity Corpus for Slavic Languages (2024.lrec-main)

Copied to clipboard

Challenge: This work presents a corpus manually annotated with named entities for six Slavic languages .
Approach: They propose to manually annotate a corpus of names for six Slavic languages . they use a transformer-based neural network architecture to train multilingual models .
Outcome: The corpus consists of 5,017 documents on seven topics . each entity is described by a category, a lemma, and a unique cross-lingual identifier.
Applying Natural Annotation and Curriculum Learning to Named Entity Recognition for Under-Resourced Languages (2022.coling-1)

Copied to clipboard

Challenge: Existing approaches to build NLP models for low-resourced languages rely on machine translation or cross-lingual transfer.
Approach: They propose to use natural annotations to build synthetic training sets from resources not originally designed for the target downstream task.
Outcome: The proposed model achieves the F1 score of 0.78 for Belarusian starting from zero resources compared to the baseline of 0.63 for English . the proposed model can be fine-tuned to reflect linguistic properties, such as the grammatical case and gender, for the Slavic languages.
A Closer Look on Unsupervised Cross-lingual Word Embeddings Mapping (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods for word embeddings are limited to a single, unannotated corpus, which means that word representations with similar meaning in distinct languages can be very different.
Approach: They propose an unsupervised method for cross-lingual word embedding mapping that uses stochastic initialization and isometric initialization to verify the method's robustness.
Outcome: The proposed method is robust on different embedding representations and new language pairs, particularly those involving Slavic languages like Polish or Czech.
Modeling the Impact of Syntactic Distance and Surprisal on Cross-Slavic Text Comprehension (2022.lrec-1)

Copied to clipboard

Challenge: Using symmetric measures of insertion, deletion and movement of syntactic units, we investigate phonetic and orthographic asymmetries between selected languages.
Approach: They focus on the syntactic variation and measure syntaktic distances between nine Slavic languages using symmetric measures of insertion, deletion and movement of syntak units in parallel sentences of the fable “The North Wind and the Sun”.
Outcome: The proposed measures are validated on spoken and written cloze tests for Slavic native speakers to determine whether variations in syntax lead to slower or impeded intercomprehension of Slav texts.
nEMO: Dataset of Emotional Speech in Polish (2024.lrec-main)

Copied to clipboard

Challenge: Existing datasets covering Slavic languages do not accurately represent basic emotional states.
Approach: They propose to use a Polish corpus of emotional speech to represent basic emotional states.
Outcome: The proposed corpus represents six emotional states in Polish, with 9 actors participating in the study.
Polish-ASTE: Aspect-Sentiment Triplet Extraction Datasets for Polish (2024.lrec-main)

Copied to clipboard

Challenge: Aspect-Sentiment Triplet Extraction (ASTE) is one of the most challenging and complex tasks in sentiment analysis.
Approach: They propose to use customer opinions of hotels and purchased products in Polish to extract ASTE triplets that contain an aspect, its associated sentiment polarity, and an opinion phrase that serves as a rationale for the assigned polarities.
Outcome: The proposed datasets contain customer opinions about hotels and purchased products expressed in Polish and are available under a permissive licence and have the same file format as the English datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations